Papers with data-driven methods
EVIDENCEMINER: Textual Evidence Discovery for Life Sciences (2020.acl-demos)
Copied to clipboard
Xuan Wang, Yingjun Guan, Weili Liu, Aabhas Chauhan, Enyi Jiang, Qi Li, David Liem, Dibakar Sigdel, John Caufield, Peipei Ping, Jiawei Han
| Challenge: | EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences. |
| Approach: | They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
| Outcome: | EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
Towards Unified Representations of Knowledge Graph and Expert Rules for Machine Learning and Reasoning (2022.aacl-main)
Copied to clipboard
| Challenge: | Empirical study shows superiority of proposed method over time-tested knowledge-driven and data-driven methods. |
| Approach: | They propose a cognitive knowledge graph that unifies expert rules and relational facts as the substrate of machine learning and reasoning models. |
| Outcome: | Empirical results show the proposed method superior to time-tested methods . the proposed model can perform both learning and reasoning with labeled data . |
How Do Large Language Models Perform in Dynamical System Modeling (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent data-driven methods often use graph neural networks (GNNs) to learn interactions between objects. |
| Approach: | They propose prompting techniques for dynamical system modeling and evaluate their performance . they find that large language models demonstrate competitive performance without training . |
| Outcome: | The proposed methods show competitive performance without training compared to state-of-the-art methods in dynamical system modeling. |
XferBench: a Data-Driven Benchmark for Emergent Language (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods to teach models to "language" are full of bias, toxicity, and potential intellectual property violations. |
| Approach: | They propose a benchmark for evaluating the overall quality of emergent languages using data-driven methods. |
| Outcome: | The proposed benchmark is based on utterances from the emergent language and is validated using human, synthetic, and emergentic language baselines. |
Social Norms-Grounded Machine Ethics in Complex Narrative Situation (2022.coling-1)
Copied to clipboard
| Challenge: | Recent studies focus on data-driven methods to judge the ethics of complex real-world narratives but face two major challenges: they cannot handle dilemma situations due to a lack of basic knowledge about social norms; and they focus on sparse situation-level judgment regardless of the social norm. |
| Approach: | They propose to complement a complex situation with grounded social norms by a norm-supported ethical judgment model in line with neural module networks to alleviate dilemma situations and improve norm-level explainability. |
| Outcome: | The proposed model improves state-of-the-art performance on two narrative judgment benchmarks. |
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)
Copied to clipboard
Piroska Lendvai, Maarten van Gompel, Anna Jouravel, Elena Renje, Uwe Reichel, Achim Rabus, Eckhart Arnold
| Challenge: | a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well. |
| Approach: | They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data . |
| Outcome: | The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP . |
Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing factor mining models are inefficient and inefficient, resulting in a significant challenge to extract interpretable factors. |
| Approach: | They propose a model that integrates the strengths of both neural and symbolic models for factor mining. |
| Outcome: | The proposed model surpasses the SOTA RankIC and RankICIR in predicting S&P 500 returns on real-world stock market data. |
Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition (L18-1)
Copied to clipboard
| Challenge: | a new general framework for sign recognition from monocular video is presented . the framework exploits state-of-the-art learning methods while incorporating features based on what we know about the linguistic composition of lexical signs. |
| Approach: | They propose a general framework for sign recognition from monocular video . they exploit state-of-the-art learning methods while incorporating features from linguistic information . |
| Outcome: | The proposed framework exploits state-of-the-art learning methods while incorporating features based on what we know about linguistic composition of lexical signs. |
Syntax-Aware Opinion Role Labeling with Dependency Graph Convolutional Networks (2020.acl-main)
Copied to clipboard
| Challenge: | Opinion role labeling (ORL) is a fine-grained opinion analysis task . due to the scarcity of labeled data, ORL remains challenging for data-driven methods due to its complexity and complexity. |
| Approach: | They propose to integrate syntactic knowledge into ORL models by comparing and integrating different representations and using dependency graph convolutional networks to encode parser information at different processing levels. |
| Outcome: | The proposed model achieves 4.34 higher F1 score than the current state-of-the-art. |
Identifying Aspects in Peer Reviews (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to peer review are limited in how they identify aspects . a growing volume of peer review submissions is straining the process . |
| Approach: | They propose a data-driven schema for deriving aspects from peer reviews . they propose augmented peer reviews and show how it can be used for community-level review analysis. |
| Outcome: | The proposed approach can be used to support peer review, but lacks formal definition of aspect . it also shows that the choice of aspects can impact downstream applications . |
A Large-Scale Corpus for Conversation Disentanglement (P19-1)
Copied to clipboard
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros C Polymenakos, Walter Lasecki
| Challenge: | a dataset of 77,563 messages manually annotated with reply-structure graphs disentangles conversations and defines internal conversation structure. |
| Approach: | They use a dataset of 77,563 messages manually annotated with reply-structure graphs to disentangle conversations and define internal conversation structure. |
| Outcome: | The new dataset is 16 times larger than all previous datasets combined and includes adjudication of annotation disagreements and context. |
Understanding the Language of Political Agreement and Disagreement in Legislative Texts (2020.acl-main)
Copied to clipboard
| Challenge: | Despite the fact that state-level legislation is rarely discussed, it has a dramatic influence on the everyday life of residents of the respective states. |
| Approach: | They propose a large-scale dataset linking state bills and legislator information, geographical information about their districts, and donations and donors’ information. |
| Outcome: | The proposed model improves over strong text-based models by integrating the state-level text and the legislative context. |
How Does the Experimental Setting Affect the Conclusions of Neural Encoding Models? (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have shown that neural encoding models explore brain language processing using naturalistic stimuli. |
| Approach: | They propose a block-wise cross-validation training method and an adequate data size for increasing the performance of neural encoding models. |
| Outcome: | The proposed training method and data size can significantly decrease the performance of neural encoding models in the temporal and frontal lobes. |
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent large-scale neural response generation models (RGMs) have made significant progress but still struggle to generate semantically appropriate responses. |
| Approach: | They build a large dataset of model-generated contradictions for the first time and analyze the results to gain valuable insights into their characteristics. |
| Outcome: | The proposed dataset significantly improves the performance of data-driven contradiction suppression methods. |
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct (2025.findings-acl)
Copied to clipboard
Run Luo, Haonan Zhang, Longze Chen, Ting-En Lin, Xiong Liu, Yuchuan Wu, Min Yang, Yongbin Li, Minzheng Wang, Pengpeng Zeng, Lianli Gao, Heng Tao Shen, Yunshui Li, Hamid Alinejad-Rokny, Xiaobo Xia, Jingkuan Song, Fei Huang
| Challenge: | a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling . |
| Approach: | They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution. |
| Outcome: | The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data. |